Skip to content

removing hardcoded path - #40

Open
084intheImpala wants to merge 35 commits into
trojsten:masterfrom
084intheImpala:master
Open

084intheImpala wants to merge 35 commits into
trojsten:masterfrom
084intheImpala:master

Conversation

@084intheImpala

@084intheImpala 084intheImpala commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

fixing some spacing issues after punctuation

sesquideus and others added 30 commits September 30, 2026 13:19
Eight `core/i18n/*.yaml` carried the note "anything absent falls back to English". **No such
fallback has ever existed** -- `LocalisedWords.__getitem__` boxes the word and records it, and
`default.yaml` deliberately holds no `words:` precisely so that `merge()` cannot invent one.
The note described code that was never written, and it is the dangerous direction to be wrong
in: a translator reading it would believe that omitting a word is safe. `en.yaml` said the
opposite in the same breath. All eight now say what the code does.

The real gap was the other half. A miss was a `log.warning` and the process exited 0, which is
not a safety net -- it is exactly how volume 19 printed `Missing file …onion…!` on page 42 in
every language for years while `make` stayed green. `JinjaCLIInterface.run` now logs an error
and raises `MissingWordsError`.

**After the output is written, not instead of it.** The booklet still builds with every hole
boxed in red, so a translator sees all the gaps in one pass rather than one per rebuild, and
`LocalisedWords` stays lazy so a Polish build does not fail over a word Polish never asks for --
`21/troll-science` writes translated subscripts in four of its six languages and differently in
the other two, and still renders clean in all of them. Only then does the render exit nonzero.

Safe today: `core/audit`'s `word-missing` reports nothing anywhere under `source/`, and all
5473 real markdown files in `phys` render clean in their own languages, missing-word failures
zero. The change bites only when someone introduces a gap, which is when it should.

Three tests, both halves as the audit convention requires: that a miss raises, that the output
is on disk with the box before it does, and that a render asking for nothing missing still
succeeds -- the quiet case, which is the one that would catch this being over-eager.

`core/audit/checks.py` and CLAUDE.md both described the old behaviour and now describe this;
the audit check keeps its job, which is finding the gap in every language at once without a
build having to be attempted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three problems carried this tag and lost it -- `26/hairy` (a hair-growth rate into a day),
`27/takeoff` (m/s^2 into km/h^2) and `25/ascent` (an ascent minus an elevation). Each is a real
calculation, however short, and nothing in any of them collapses.

The one-line definition admitted both readings, and a sweep over the whole archive found it
being used for both: *trivially easy, dressed up as hard* and *a scary statement that boils down
to nothing*. Only the second is meant. The test is now stated -- a given goes unused, or the
answer refuses the question, or it is a constant the statement already handed you -- with the
three removals as the worked examples, so the question does not come up a fourth time.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Czech, Slovak and Portuguese all repeat the hyphen when a word breaks at one: `anti-` then
`-inflamatório`, never `anti-` then `inflamatório`. ČSN and STN 01 6910 state it for the
spojovník; Portuguese needs it most, because enclitic pronouns put a hyphen inside ordinary
verbs -- `encontra-se`, `deu-lhe`, `colocou-o`, all of which are in volume 28 already.

`core/filters/hyphens.lua` puts `\rephyphen` where a declaring language has a hyphen between
two letters, the `spacing.lua` pattern and for the same reason: a hyphen in these sources is
far more often *not* prose -- `northern-sun.svg`, `{#fig:drag-queen}`,
`forbid-literal-units=false`, `$a-b$` -- and pandoc has already sorted those into node types a
`Str` rule cannot reach. Opt in with `typography: repeat_hyphen: true`; a language that
declares nothing keeps the plain hyphen. Unlike `spacing.lua` it is LaTeX only: no Unicode
character repeats a hyphen at a break, U+00AD inserts one without repeating it.

**The macro is `\mbox{-}\discretionary{}{-}{}` and neither half is arbitrary.** Both cost a
rebuild to find, and babel's own answer fails on the first:

- The pre-break list is **empty**, the hyphen written before the discretionary. TeX governs a
  discretionary break by \exhyphenpenalty when the pre-break list is empty and by
  \hyphenpenalty when it is not -- so `\babelhyphen{repeat}`, which is `\discretionary{-}{-}{-}`,
  is governed by \hyphenpenalty and does not break at all in a class that discourages ordinary
  hyphenation. This way round the two penalties stay independent.
- The hyphen is in an **\mbox**, because TeX inserts a breakpoint of its own after an explicit
  hyphen, it sits *before* the discretionary, and it costs the same \exhyphenpenalty. With both
  legal TeX takes the earlier and the hyphen is not repeated. It shows only on words short
  enough for the two to compete, which is why `anti-inflamatório` looked right in testing while
  `encontraram-se` came out as `encontraram-` / `se`. `\char`\-` does not help.

The rule is inert unless \exhyphenpenalty is low enough to allow the break, so cs, sk and pt
carry `latex: exhyphenpenalty: 50`, TeX's own default. That is set per language rather than in
`dgs.cls`, whose two penalty lines are being changed by someone else right now; this works
whatever they settle on.

Verified by building, not by reading: all three solutions booklets for volume 28 build clean
and come out with **zero** hyphen-broken lines, so lowering the penalty from the ban changed
nothing that was already right -- it buys a breakpoint for lines that need one. The repetition
itself is verified under the real class at a width that forces the break.

Eight tests: that the three get the macro and `en`, `de`, `pl`, `ru` do not; that it is braced;
that maths, code, figure paths, crossref labels and numeric ranges are untouched; that a hyphen
needs a letter on each side; that HTML keeps the plain hyphen; and that the three lower
\exhyphenpenalty, without which the rest is decoration.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…r nouns

Swept all twelve languages against each one's own authority rather than assuming, because the
answer is not guessable: the two Slavic languages that repeat and the two that do not sit side
by side, and the first cut of this shipped with Polish missing.

Repeats, and now declares it:

  Czech       UJC, Internetova jazykova prirucka 164, after CSN 01 6910 -- already had it
  Slovak      STN 01 6910 -- already had it
  Portuguese  Acordo Ortografico 1990, Base XX item 6 -- already had it
  Polish      PWN rule [196]: czarno- / -biale, warminsko- / -mazurskie     -- ADDED
  Spanish     RAE, Ortografia 2010: lexico- / -semantico                    -- ADDED

Does not, and is left alone: German (Duden -- the Bindestrich "gilt bei der Worttrennung am
Zeilenende gleichzeitig auch als Trennungsstrich"), French (OQLF -- "on ne le repete donc pas au
debut de la ligne suivante"), Hungarian (AkH. 238 -- only "szakmunkakban"), Russian (Milchin --
"обычно не повторяется", and some advise never breaking at a hyphen), Ukrainian (the Pravopys
does not state it), English (no style guide prescribes it), Farsi (hyphens are rare; compounds
join with ZWNJ).

**Spanish carries an exception the other four do not, and it is the one that matters here.**
RAE: "Excepto cuando la palabra que sigue es un nombre propio que empieza con mayuscula" --
the capital already shows the hyphen is not a division mark, so `Ruiz-` / `Gimenez`. Physics is
made of these: Gay-Lussac, Navier-Stokes, Bose-Einstein, Gutenberg-Richter. Slovak is the exact
opposite, STN giving `Rakusko-Uhorsko` and `Bratislava-Ruzinov` as cases that do repeat, so the
exception cannot be generalised. `typography: repeat_hyphen_not_before_capital` carries it and
only `es` sets it.

Checked on real source text rather than on examples: `26/earthquakes` names the Gutenberg-Richter
law in six languages, and it now converts to `Gutenberga\rephyphen{}Richtera` in Polish and
`Gutenbergovym\rephyphen{}Richterovym` in Czech while Spanish keeps the plain `Gutenberg-Richter`.
Volume 28 builds clean in both new languages with zero hyphen-broken lines, so nothing that was
already right moved.

Eleven tests now: that all five get it and the seven that do not are untouched, that Spanish
exempts a following capital and that the other four do not, and that every repeating language
also lowers \exhyphenpenalty, without which the rest is decoration. CLAUDE.md carries the table
with a quotation from each authority, so the next person does not have to sweep it again.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Russian prompted this and then turned out not to need it. Checking whether Russian wants a
blanket \exhyphenpenalty=10000 found the opposite: its rules *permit* the break -- `сет-клин`
and `монотип-клавиатура` are their own examples -- and merely do not repeat the hyphen, so a ban
would forbid what the norm allows. Russian stays out of the table.

What Russian does forbid is `2-местный`, `Боинг-767` and multi-digit numbers, and that is not
Russian-specific: Czech typography forbids dividing `3-dílný` or `10-procentní` at the
spojovník, and the Ukrainian Pravopys forbids separating a grammatical ending from its digits
(`10-й`).

**That is a regression the previous two commits introduced.** The filter skipped digit-adjacent
hyphens, which left them as plain hyphens -- and the five repeating languages now carry
\exhyphenpenalty=50, so a hyphen merely left alone is freely breakable where before it was
discouraged at 100 or banned at 10000. Leaving a case out is not the same as handling it.

So the filter now says it positively: a digit on either side emits `\nbhyphen`, an `\mbox{-}`
with no discretionary. The \mbox is doing the same job as in `\rephyphen` -- hiding the
character from TeX's explicit-hyphen rule, which would otherwise offer a break there whatever
the discretionary does.

Real: 70 in the Slovak sources (`10-stupňovej`, `12-krát`, `1-2`) and 4 in Czech (`207-krát`,
`70-kilový`). Volume 28's `balance-me` has `8-g`, which now converts to `8\nbhyphen{}gramov` in
Slovak and stays `8-g` in English; both volumes rebuild clean. Verified under the real class at
a width that forces a break: `10-stupnovej` and `123-456` move whole to the next line while
`cesko-slovensky` still splits and repeats.

Two tests, both halves: that all five languages make it unbreakable and emit no `\rephyphen`
there, and that `en`, `de` and `ru` are untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The rule was being kept by accident. Statement displays were written by hand as `$$ … $$`,
which produces no label -- and that is exactly why none of them had ever been hoisted: `|disp`
attaches `{#eq:<pid>:<key>}` unconditionally, so hoisting a statement would have started
numbering it. Hand-written and drifting, or hoisted and numbered; there was no third option.

`MathObject` now takes a `labelled` flag and `JinjaCLIInterface.build_context` sets it from the
file being rendered: `problem.md` and `problem-extra.md` get no label, everything else does.

**By file and not by filter**, deliberately. A `|dispu` beside `|disp` would have to be
remembered in nine translations of one statement and got right every time; this cannot be got
wrong. It is also what lets one `eq:` entry serve both halves of a problem, which is the whole
reason a statement can be hoisted at all -- `26/earthquakes` states the Gutenberg-Richter law
and opens its solution with it, and the same `(§ eq.grl|disp(',') §)` is now unlabelled in the
statement and labelled in the solution. Labelling both would have been a duplicate label.

All three display filters attach it, so all three drop it: `disp`, `align` and `arr`.

Nothing in the repository moves. There are zero labelled displays in any `problem.md` anywhere
-- the rule held before, by hand -- so this only changes what happens when a statement *is*
hoisted, which is what the next commit does.

Four tests: that a statement carries no label and a solution does, that one entry used in both
yields exactly one `{#eq:}`, and that `align` and `arr` obey it too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`problem.md` beside `solution.md` is Náboj's taxonomy, and core had no business knowing it.
Seminar nests `problem.md` five levels deep under competition/volume/semester/round; scholar
also has `text.md`, a lecture, whose equations *should* be numbered. A `startswith('problem')`
in `core/builder/renderer.py` was one module's convention imposed on the other two -- the same
mistake `editor.yaml` exists to avoid, and the one the audit records as `audit: true`.

So `CLIInterface.unnumbered_files` is a class attribute, empty in core, and
`modules/naboj/builder/renderer.py` names its own two files. A module that says nothing numbers
everything, which is what all three did before this existed. Nothing outside Náboj changes --
all eleven statement displays in the repository are Náboj's, seminar and scholar have none --
but the day seminar adds one it will not silently inherit a rule nobody wrote for it.

Also records why there is no explicit `label=` on `|disp`, since the alternative is the obvious
one to reach for. **The call is not hoisted even when the equation is**: 5264 call sites name
2579 distinct equations, so more than half of all equations are written out once per language.
`label=False` would be written six times for `26/earthquakes` and could be forgotten in one,
giving an equation numbered in Slovak and not in English -- the one-language drift hoisting
exists to prevent, in a new costume.

And those argument lists already vary per language on purpose: 141 equations are called with
different terminal punctuation in different translations, because the punctuation follows the
sentence around the display and that is the translator's call. A `label=` would sit in exactly
that argument list and be varied by the same hand. This also kills the check proposed earlier
for catching such disagreements: measured against the sources it is 141 false positives and 0
true ones.

A per-entry flag in `meta.yaml` would be un-driftable but cannot express the case that made the
hoist possible at all: the same entry unlabelled in the statement and labelled in the solution.

Three tests: that core holds no list while Náboj holds its two, that `problem-extra.md` is a
statement too, and that an answer file is not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`unnumbered_files` was a list of exceptions: it said where a module differs from a default and
left the rest to be inferred. `equation_numbering` is a mapping instead, so reading a module's
declaration gives its whole taxonomy:

    'problem.md':       False      'solution.md':        True
    'problem-extra.md': False      'answer.md':          True
                                   'answer-extra.md':    True
                                   'answer-also.md':     True
                                   'answer-interval.md': True

The answer files are listed although none of them carries a display anywhere in the repository.
They are part of Náboj's taxonomy, so the decision is recorded rather than inherited, and a new
test pins the keys to the two rule families in `module.mk` -- the way `test_audit` already pins
`TRANSLATED_FILES` -- so a file added to the module cannot pick up a default unnoticed.

**A file the mapping does not mention keeps its numbers**, and that default stays implicit on
purpose. It is the conservative direction: a module that declares nothing is unaffected, which
is every module but Náboj today, and a file type added later is numbered until somebody decides
otherwise rather than silently losing its labels. Losing a label is the failure that is hard to
see -- a missing equation number looks like a design choice, where an unexpected one does not.

Core still holds no opinion. Behaviour is unchanged: `problem.md` and `problem-extra.md` are
exactly the two that were unnumbered before, and volume 26's Slovak booklet rebuilds identically.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A picture that shows a number had to have that number typed into it by hand, so it drifted from
the `meta.yaml` that computes the same number for the prose. `pool/filling-large/solved.tikz`
carries a comment saying exactly that. Pictures now take the route `.gp` has always had,
`source/ -> render/ -> build/`, with the `meta.yaml` beside them as context.

The Makefile already half-believed this worked: its `pathlang` comment opens with "Language a
picture (`.tikz`, `.gp`) is rendered in ... it selects the locale the Jinja tags format numbers
with", while `.tikz` went straight from `source/` to `build/` and a tag in one reached LaTeX
verbatim. `standalone.jtex` splices the drawing in as `(* content *)`, and Jinja substitutes a
variable's value literally rather than rendering it again.

**The block and comment tags had to move, because the defaults are the formats' own notation.**
`(#` is how SVG refers to anything it defines -- `url(#linearGradient42)`, `url(#Arrow1Mend)`,
every clip path, mask and filter -- 11081 of them across 2006 files, 792 of those files in a
module that builds; Jinja opens a comment and never finds `#)`. `(@` is chemfig's arrow between
named nodes, `\arrow(@c2--c4){0}[-90]`, in two live chemistry pictures. Neither can be spelled
around, so `PictureJinjaRenderer` moves the tags to `(@§ … §@)` and `(#§ … §#)` and leaves the
variable tag alone, `§` being the one character absent from every picture in the tree. `§` hugs
the content in all three and the character outside says which tag it is; `(§@` reads better and
does not work, because Jinja matches the variable tag first.

`.gp` stays in the Markdown dialect deliberately -- gnuplot comments with `#` but never writes
`(#` or `(@`, so it has nothing to dodge and its existing templates would be rewritten for
nothing. `PICTURE_SUFFIXES` is the whole rule and `renderer_for` is tested both ways.

The web targets move too, or `output/` would ship the tags.

**The conversion had to be a no-op and was.** Nothing in the repository carried a tag: of 1070
live pictures rendered, 1059 came back byte-identical and 11 differed only by a stripped
trailing blank line, which is inert in TikZ and SVG alike. `phys/29` builds to 44 and 50 pages
as before, `chem/03` to 42, and a seminar picture through both the PDF and the web leg.

An Inkscape 1.4 round trip on a real drawing confirms a tag survives a save inside its `<tspan>`
and still renders, which is the claim the SVG half rests on.

Two things found on the way, neither fixed here. `seminar/FKS/40/2/3/06/meta.yaml` writes
`value:` where the schema wants `magnitude:`, so it has never loaded and its `problem.md` fails
the same way -- the picture simply fails with it now. And seminar's and scholar's `.gp` rules
pass `$(lang)` where naboj passes `pathlang`; the new rules use `pathlang` and the old ones are
left for their own commit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A drawing is a template now, and a tag in one cannot be `(§ v §)`: that emits
`\qty{3}{\metre\per\second}`, and `rsvg-convert` sets SVG text with a system font, so the
backslashes go on the page. The workaround was `(§ v.mag §) m/s` with the unit left as a
literal beside it -- which puts half the value back in the drawing, where it drifts, which is
the thing templating a picture was for.

`(§ v|txt §)` writes `3 m/s`. `txtg` is the `g` variant, for anything `txt` would write as
`0.0000000000…`, and both carry the `txt0`-`txt9` family.

The unit is **pint's own `~P`** rather than a table here: short symbols, Unicode superscripts,
`/` and `⋅`. A table of ours would have to carry every unit the sources use and would drift
from the registry the values are built in; pint's comes out of that registry, so a unit that
can be computed can be printed. The magnitude is ours, through the same `format_float` the `|f`
family uses -- pint would print `3.0 m/s` where every other filter in the repository prints `3`,
and a drawing sits beside prose that used one of them. An exponent is superscripted to match the
unit's powers, `6.674e-11` becoming `6.674×10⁻¹¹`; `cut_extra_one` having already dropped the
mantissa, `1e15` gives `10¹⁵` with no stray multiplier.

A bare degree closes up, `45°`, mirroring the `\ang` case in `PhysicsQuantity.__format__` and
for the same reason; a rate of degrees does not, so `30 deg/s`. `°C` keeps its space.

Two things it inherits rather than fixes, both stated in the docs: pint's denominator order,
so water's heat capacity is `4180 J/K/kg`; and the unit the arithmetic produced, so `omega * R`
is `3 m⋅rad/s` and wants `.to('m/s')` -- which is what `29/curveball`'s `eq:` entries already
write.

A range, a list and a product have no plain spelling yet and say so by name.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`75 cm – 77 cm`, `1 m, 2 m, 3 m`, `3 cm × 4 cm × 5 cm`. The separators are read off
`core/latex/siunitx.tex` rather than chosen here, because the filter's job is to print what the
booklet prints: `range-phrase = {\text{ -- }}` is an en dash with spaces, `list-separator` is a
comma and a space, and `range-units` and `list-units` are both `repeat`, so each part carries the
unit. `QuantityList.__init__` has already coerced every element to a common unit, so repeating it
is just rendering each one.

**A range goes through `QuantityRange._outward`, not through its two endpoints.** A range here is
the set of answers a marker accepts: rounding each end to nearest shrinks it and turns away
correct work, which is the whole argument in `_outward` and the reason `29/bouncy-v` printed
`3.7 - 3.8` and rejected the solver who used the exact `g`. Formatting the endpoints directly
would have put that back, quietly, in the one place nobody would look for it -- a figure.

So `txt` and `__format__` round identically, and the test pins them against each other rather
than against a literal: `text(band, 1)` is `3.6 m – 3.8 m` while `format(band, '.1f')` is
`\qtyrange{3.6}{3.8}{\metre}`, and if either grid moves the test fails. `29/folding-bath`'s real
interval agrees the same way, `83.5 l – 84.0 l` against `\qtyrange{83.5}{84.0}{\litre}`.

A `MathObject` still raises: it is LaTeX by nature and has no plain spelling.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…a label

Outward rounding exists for one reason: `answer-interval.md` is the set of answers a marker
accepts, so rounding its ends to nearest shrinks that set and turns away correct work. That is
the `29/bouncy-v` defect and it is why `__format__` floors the minimum and ceils the maximum.

Nothing `|txt` prints is an answer. A range in a drawing labels a span -- an axis, a dimension --
and moving its ends outward would state a width nobody measured, for a reason that does not
reach there. So each end is simply formatted, and the two paths now disagree: a band of
[3.67749, 3.75] at one decimal is `3.7 m – 3.8 m` here and `\qtyrange{3.6}{3.8}{\metre}` there.

The previous commit had this backwards, reasoning from the mechanism rather than from what the
number is for.

Both behaviours are asserted side by side so neither can be quietly turned into the other, and a
band already on its grid is pinned too, since that is the case where they agree and the
difference could otherwise go unnoticed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`\Implies`, `\Iff`, `\ImpliedBy`, `\LAnd`, `\LOr`, `\LXor`, `\LNand` have been in
`core/latex/math.tex` for some time and CLAUDE.md never mentioned them, so the sources wrote
the expansion out instead: 61 sites in `seminar`, at two different widths because nothing said
which was meant.

Also records the case that is *not* one -- a bare `\Rightarrow` opening a row inside an
`aligned`, where the column already supplies the spacing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`core/builder/jinja.py` registers 182 filters and 32 globals on the Markdown
environment. 57 of those names appear anywhere in `source/`; the other 125 are
unreachable in practice, because the only way to find one was to read the table
in `jinja.py`. This writes them all down.

`core/builder/reference.py` is that table as data -- 146 entries covering every
filter root, every global, the quantity attributes, the `%` range operator, the
`const` and `i18n` namespaces, and the `meta.yaml` keys at all five levels of the
hierarchy. `docs/filters.md` is generated from it by `python -m
core.builder.reference`, and `--check` exits nonzero when the file on disk is
stale, which is the CI shape `core/audit/macros.py` already has.

Not a scrape of the docstrings, deliberately. Those are written for whoever
maintains the code -- `num`'s explains why it is idempotent, `_unop`'s explains
contamination -- and ten filter roots have none at all. An author needs "what do
I type and what comes out", which is a different text about the same function.

What makes it trustworthy is `core/tests/test_reference.py`, and the second test
is the point of the whole thing:

- the table's names must equal the registered names **in both directions**, so a
  filter added without an entry fails the suite and an entry naming a filter that
  was deleted fails too;
- **every example is rendered and compared** with the output printed beside it.
  All 102 of them. The expected outputs were produced by running the code, so a
  filter whose output moves takes this file red with it rather than quietly making
  102 printed examples wrong;
- `docs/filters.md` must match the generator, so it cannot be hand-edited into
  something the table does not say.

Each was verified by mutation: deleting an entry, breaking one expected output,
hand-editing a line of the document and adding an undocumented filter to
`jinja.py` each fail, and the clean tree passes.

Two things the inventory turned up that were not written down anywhere:

- **the two environments are disjoint.** `MarkdownJinjaRenderer` serves `.md`
  (and `.svg`/`.tikz` through `PictureJinjaRenderer`); `StaticRenderer` serves
  `.jtex`. `|disp` is written 4760 times and exists only in the first, `|upnth`
  11 times and only in the second, and either mistake is found by a failed build
  and nothing else. A test asserts the two filter tables do not intersect;
- **exactly one filter shadows a Jinja builtin**: `e`, which overwrites `escape`.
  Harmless, because the environment sets `autoescape=False` and nothing writes
  `|e` -- but pinned, so a second one cannot be added without the decision being
  made on purpose. Shadowing `|join` or `|round` would change what existing
  sources mean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A `(§ … §)` was one token, so the pane could colour it and say nothing about what
was inside it. It is now emitted as a run of siblings -- the filter after a `|`,
the global before a `(`, the attribute after a `.`, the `%` range operator -- each
carrying `data-ref="<kind>:<name>"`, keyed exactly as `/api/reference` keys its
entries so the hover does one dictionary lookup and no parsing.

**Flat siblings, not nested spans.** Nothing new was needed for the split: giving
the inner names a higher priority than the tag rule makes the existing carving do
it, the same mechanism that already lets a tag win against the `$…$` it sits
inside. Keeping the output flat also keeps `highlight`'s shape, so the two
read-only panes that call it are unaffected and the tests that parse its spans
keep working.

The meta pane gets the same treatment, and it matters more than it looks: 158 of
the 160 `PQ` in the repository are in a `derived:` entry and two are in a tag, and
`QR` is never written by name at all -- an interval is `(result % result_approx)`,
85 files of it. A hover confined to Markdown would have missed nearly all of it.

Four distinctions the rules have to make, each with a test for the quiet case:

- `|snap(0.5)` is a filter, not a call, and `d.to('cm')` is an attribute, not a
  global. Both match two rules; `INNER`'s order decides. That order is now applied
  in `collectReferences` rather than only in `highlight`, because the F1 lookup
  reads the list directly and was seeing `.to` as a global as well -- two paths,
  one of them wrong;
- `eq.kin` is an equation and `const.g` is a constant. Neither is an attribute
  *of* anything, so the name after a namespace stays unclaimed rather than
  offering `.kin` as though `PhysicsQuantity` had one;
- `'%s-%02d'` is a format string. The range operator needs whitespace on both
  sides, which is what excludes it;
- `values:` has two levels and only the deeper one is the table's. `h:` is the
  author's own name for a quantity and there is nothing to look up. The name level
  is read off the first indent rather than assumed to be two spaces.

Swept over all 11421 `.md`, `.yaml` and `.gp` files in every module: **nothing
loses a character**, and 30609 references resolve against the table. The rest are
seminar's and scholar's own meta schemas, which this does not document; a miss
shows nothing, which is the right outcome.

Also fixed a latent trap in the test helper: its span pattern did not allow for an
attribute, so a span carrying `data-ref` fell through to the *next* one and the
body of one token swallowed the tokens between -- a corrupt reading that no
assertion would have looked wrong.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`/api/reference` serves the same table `docs/filters.md` is generated from, so the
tooltip and the document cannot disagree and both are held to the code by
`core/tests/test_reference.py`. Sent whole and once: it never changes while the
server is up, and the alternative is a request per hover.

**The hit test had to be worked around.** The overlay is `pointer-events: none`
under a textarea that is not -- which is the only way the two layers work at all,
since the text you see is the overlay's and the text you edit is the textarea's --
so `elementFromPoint` returns the textarea and never a token. `tokenAtPoint`
inverts the two for the duration of one hit test and restores both before the
handler returns. Verified in Firefox rather than argued: a probe at the centre of
each token resolves it, a point in the prose resolves nothing, and the textarea
still owns the pointer afterwards.

The alternative -- walking every `[data-ref]` span and asking each for its rects
-- is O(tokens) forced layouts per mouse move, and a solution with four hundred
tags would make the pane crawl.

F1 answers the same question of the caret, and it calls `collectReferences`
directly: no DOM, and the keyboard and the mouse resolve out of one definition of
where a name is rather than two.

Smaller things the rendering needed:

- the tooltip hangs off `body`, not `.code-editor`. That is the only positioned
  ancestor but it is also `overflow: hidden`, and a token at the end of a line is
  exactly where a clipped tooltip would be;
- it is opaque. At 0.97 the code underneath still shows through the prose;
- a tag is a run of spans now, so the wash is shared across all of them and the
  2px radius is gone -- rounding each piece draws the seams the split is meant to
  be invisible across;
- `[data-ref]` gets a dotted underline, the same affordance the audit page uses
  for a word that has an explanation. The `cursor: help` cannot show, because the
  textarea owns the pointer, but the underline is what says there is anything to
  hover at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
enschema reads a **subscripted generic as a callable** and validates by calling
it. So `dict[K, V]` only ever asked whether the value could be passed to
`dict()`, and `list[V]` whether `list()` would take it — which it always will,
and `list('aiko')` is `['a','i','k','o']`. Neither the key pattern nor the value
type was ever looked at.

Twelve declarations in two schemas were written that way, and every one of them
was decoration.

**`meta.yaml`'s five content keys.** What passed, each failing later and
somewhere else: a `derived:` entry holding a number rather than an expression;
an `eq:` key that is not an identifier; a `words:` term with a bare string where
its languages belong; and a `values:` entry with a misspelt `magnitude`, which
surfaced as a bare `KeyError: 'magnitude'` out of a constructor, with nothing in
it to name the entry or the file.

A `values:` entry's own keys are now spelled out in `VALUE_ENTRY` — exactly the
keywords the constructor takes. Three decisions in it:

- **The mapping form goes last in the `Or`**, because that is the alternative
  whose error gets reported. A misspelt key now fails as `Missing key:
  'magnitude'` rather than `should be instance of 'PhysicsConstant'`, which
  names the one form nobody wrote.
- **`magnitude` may not be a string.** That is what catches the YAML trap at the
  key that caused it: `1e15`, `1e+15` and `1.0e15` all parse as *strings*, and
  the first sign of trouble used to be a `derived:` expression reporting
  `unsupported operand type(s) for /`. The *entry* may still be a bare string —
  that is the documented way to pass LaTeX through — and a bare number is a
  dimensionless given written in one line.
- **`unit:` with nothing after it stays legal**, because that is how a ratio or
  a coefficient of friction is written.

**`core/i18n/<lang>.yaml`.** The same hole, and it was hiding a schema that had
drifted: `prefixes` was declared `dict[str, dict[str, str]]` and is really keyed
by the exponent, `3: {name: kilo, symbol: k}`. Nothing had ever compared the two.
The `typography:` lists are the sharp end — `singles:` written as a bare string
would have validated and been split into its letters, and `spacing.lua` would
then have glued every one of them to the word after it.

Verified before landing, because tightening a schema can only be trusted that
way: **every `meta.yaml` under `source/` — 2524 files, 961 of them with content
keys — validates**, as do all 13 locales, and a full audit over 33 scopes reports
no `meta-invalid`. 25 tests, each in the two halves CLAUDE.md asks for: the thing
that is wrong is rejected, and the thing that merely looks like it is not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The hints were written with the worked examples and the counts beside them — a
filter annotated "right for 297 of the 373 such blocks", a key illustrated by
three problem ids, an operator introduced as "85 files of". That reads as notes
about this tree rather than documentation of what a tag does, and it goes stale
the moment a volume is added. Twenty-one summaries and notes now say what the
thing is and when to reach for it, and nothing about where it happens to be
used; the generated document carries no problem id at all.

The physics that justified a rule is kept wherever it is the explanation —
`digits` still says why propagating it would claim a precision nobody measured,
and `prints_exactly` still says why the round-trip allows a thousand ulps — but
stated as the general case rather than as one problem's history.

One of those notes had also been made false by the schema commit: `magnitude:`
written as `1e15` is no longer accepted-and-deferred, it is rejected by name.
The CLAUDE.md bullet that said the same thing is corrected with it, and the
schema itself gets a section there.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`.ref-example` was `pre-wrap`, so a tag broke at whichever space fell at the edge
of the box and `(§ eq.kin|disp('.') §)` read as two tags on two lines. It is
`pre` now, and the box widens for it instead — up to 54rem, which fits the
longest example in the table, 104 characters of `array` column spec. The prose
is capped at 34rem of its own so a wide example cannot stretch a paragraph to an
unreadable measure, and an inline `code` span is `nowrap` for the same reason the
examples are.

Measured rather than eyeballed: a harness renders all 146 popups and compares
`scrollWidth` with `clientWidth` on every example, signature, summary and note.
Nothing is clipped.

Anything that ever did overflow is clipped rather than given a scrollbar — the
popup is `pointer-events: none`, so a scrollbar would be one the pointer cannot
reach, which is worse than none.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`flafter` and `placeins`, and one `\FloatBarrier` at the top of `blocks/problem.jtex` and
`blocks/solution.jtex` -- which is also the end of the previous problem, so one line per block
fences every problem on both sides and `\end{document}` fences the last. Those two blocks are
used by exactly `booklet.jtex` and `solutions.jtex`, so this covers both documents.

All three are inert today, because `hacks.tex` still forces `[H]` and a float that never floats
has nothing to be barred from and nothing to rise above. They are here so that relaxing the
placement is one line rather than three, and so the guarantees are already in place when it is.

Measured rather than assumed, across six documents and 268 pages: the 28 and 29 booklets,
solutions and tearoffs come out pixel-identical, bar the two colophons, which carry a build
timestamp.

Why the barrier is the part that matters: both blocks `\setcounter{figure}{0}`, so a figure that
drifted one problem forward would print its own number under a heading whose counter has been
reset.

Two things to know before relaxing `[H]`, both measured:

- `!ht` reflows a booklet. A real float carries `\intextsep` above and below where `[H]` carries
  nothing, and in volume 28 that propagates from page 16 on -- same length, different breaks.
- The tearoff must keep `[H]` for itself and not inherit it. Its problems sit in fixed-height
  minipages, which are inner mode, where LaTeX refuses a float outright. It gets away with it now
  because a statement picture is written `![](x.svg)` with no caption and pandoc emits no float
  at all -- but 40 statement images in phys do carry a caption, 14 in volume 17 and 25 in 18.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The reference popup said what a filter *is*; it now also says what the fragment under the pointer
comes out as in this problem, in this language, on this tab -- the number, the equation, the word.

Two things are evaluable, and `evaluableSpans` in `highlight.js` is the one definition of both: a
`(§ … §)` tag, evaluated as written, and an entry's own name in the meta, for which the tag that
would reach it is synthesised -- `snell` under `eq:` answers for `(§ eq.snell §)`. Every carved
piece of a tag carries the whole tag, so pointing at the filter works like pointing at the space
beside it.

`POST /api/evaluate` does it in-process, writing nothing, and reads the *buffers* out of the
request rather than the files: the point is to answer for the text on screen, which is usually
unsaved. One request per buffer, not one per hover -- building the context evaluates every
`derived:` entry and costs the same for forty fragments as for one -- so only the first hover in
a pane waits.

It cannot disagree with the build, because `build_render_context` and `render_twice` are now the
renderer's own and both callers use them. `CLIInterface.build_context` is what is left: read the
meta, hand over the rest. `FileContext` gained `text=` so a context can be built from a buffer,
and `load_string` keeps the filename in the message, or `DuplicateKeyError` would stop saying
which file has the duplicate. A test asserts the two outputs are equal rather than trusting that
they share a function.

A meta that will not validate, or a `derived:` entry that will not evaluate, is reported in the
popup instead of a value. Strict on purpose: `make` would refuse the same meta, and a popup that
quietly accepted one it will not is the single disagreement this exists to prevent.

`test_editor_highlight`'s span regex now matches `data-` attributes by name, so a third one
cannot silently corrupt the harness the way its own comment warns about.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`submission.js` in the FKS repository is the only thing between a participant's afternoon and an
evaluator's screen, and a decoder that quietly returns the *wrong* state is worse than one that
refuses. It is pure, so QuickJS loads the shipped file directly -- the same arrangement, and for
the same reason, as `test_editor_highlight.py`.

Base64 is written out by hand in that file rather than taken from `btoa`, which would be shorter
and does exist in a browser. `btoa` does not exist in QuickJS, and a codec whose shipped path and
tested path differ is a codec with an untested path.

28 tests: every byte value and every length either side of a base64 boundary survives; a block
that has been wrapped, grouped, spaced out, hyphenated, soft-hyphenated or em-dashed by a word
processor still decodes; every one of the 2008 single-character substitutions is caught; a
truncated block, an empty one, another problem's state and a newer format are each refused with
their own name; and the character count the panel prints is the one the error message prints,
which it was not until this found it.

Skipped when `source/seminar/FKS` is not checked out, the shape `test_editor_evaluate.py` uses.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
An interactive problem is one HTML file with no dependencies -- a participant downloads it and
opens it from a folder -- so `submission.js` is inlined into each page between two markers rather
than linked. The test therefore runs the *page's* copy, and then asserts it is still character
for character the canonical `.static/web/submission.js`. Two codecs that merely both passed these
tests would still be two codecs, and the first block saved by one and loaded by the other would
be the thing that found out.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`answer.md` is what a marker compares against, so it is the value and nothing else:
`\dfrac{k(k+2)}{2k+1}`, not `\dfrac{h_2}{h_1} = \dfrac{k(k+2)}{2k+1}`. The symbol on the left
restates a question the marker already has in front of them, and costs a line of the answer
booklet per problem. In the *solution* the opposite is usually right, because the final display
is a derivation step and its left-hand side says what is being computed -- desirable there, not
required, and never carried over.

Three exceptions, and they share a reason: the left side carries information the right side
cannot. Several quantities at once, where the names are what tell them apart (seven answers).
A ratio whose members the statement does not name -- `04/energy-ratio`'s `E_k : E_p = 15 : 1`,
where a bare `15 : 1` does not say which way round it goes (three answers). And a statement that
asks for that form outright.

An `=` between two spellings of the same value is a different thing and always fine:
`\qty{960}{\giga\joule} = \qty{9.6e11}{\joule}`, or a factored and an expanded form side by
side. 152 answers do it and it is usually a kindness.

Written down rather than enforced, for now: 278 answers in phys are still of the
`symbol = value` shape, so a check would arrive with 287 findings. The archive is its own pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The fixture evaluates the inlined block on its own, where the page's own
`'use strict';` is not in front of it -- so a bare assignment to an undeclared
global passed here and threw in the browser, which is exactly what happened.
Evaluate it the way the page does, directive first in the same script, and
insist the global is reachable afterwards.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…t phys is swept

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two checks for the two shapes the phys sweep just removed, so the next one goes
red rather than accumulating:

- `value-equation-literal` -- `$X = (§ x §)$` where `values.x` declares `X`.
  Both halves of that span are the meta's: the symbol, and the relation, which
  `eq` picks by `prints_exactly` instead of taking on trust. A span printing at
  a precision is not reported: it is `|ef(N)` or `|af(N)`, and choosing between
  those is a judgement about the value.
- `value-symbol-literal` -- a bare `$X$` where `values.x` declares `X`. That is
  what makes `symbol:` worth declaring: rename it and the prose follows.

Both silent across the swept phys and across chem.

Three things they must not do, all of them met by the sweep and all of them
pinned by a quiet test. Spans come from `RE_INLINE`, never from a regex over
the raw text -- that matches across the gap between two tags and offers
`$ a $` out of `(§ q.eq §)$ a $(§ a.eq §)`, where `a` is the Slovak for *and*.
A decorated sibling (`$F_M$` beside a symbol `F`) makes the letter a family, so
a bare `$F$` is an author's call. And a symbol inside a longer span is left to
`hoistable-inline`, because the span is usually a relation and a tag inside a
literal copy of one makes three spellings out of two.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
sesquideus and others added 5 commits October 3, 2026 21:33
Two was enough for everything the sources did until a `symbol:` learned to
carry a `words:` tag. `(§ x.s §)` written *inside* an `eq:` entry is three
deep -- the entry arrives in the first pass's output, `.s` resolves in the
second and leaves the word's tag behind, and there was no third pass. The
literal `(§ words.twin §)` reached the page, in maths, where it sets as a row
of backslashes rather than failing.

`render_to_fixpoint` renders until a pass changes nothing, at least twice and
at most `MAX_RENDER_PASSES`, raising by filename rather than spinning if a
value holds a tag that resolves to itself -- which `blocks:` makes easy to
write and which used to produce silent half-expanded output.

**It costs nothing where nothing nests.** A pass over output with no tags left
changes nothing, so the loop stops after the second pass on an ordinary file,
which is exactly what it did before; only the files that need a third pay for
one. A test pins the pass count.

Three tests, and the two that matter fail under the old rule: the three-deep
symbol resolves, and the self-referential block raises instead of quietly
emitting half of itself.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The check told you to write `\begin{aligned}` into the source by hand. That
trades one local dialect for a block that is still a per-language copy, which
is the duplication the rest of this repository spends its effort removing. It
now names the `eq:` key the block's own label gives it and asks for
`(§ eq.<key>|align('.') §)`; a block that stays in the source stays as
`$${ … }$$`.

An unlabelled block gets a different message and no recommendation, because
the key becomes the label and hoisting one would give a solution's display a
number it never had. Two tests, one per branch of that.

CLAUDE.md's section says the same, and records the sweep -- 62 of 117 hoisted,
55 left -- along with the two things it took to do safely: the gate has to be
the converted TeX rather than the rendered Markdown, since the Markdown is
meant to move; and the block regex has to exclude a closer of its own, or it
starts at an unlabelled opener, closes at a later labelled one and swallows
the prose between them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`blocks/colophon.jtex` wrote "Náboj Physics international team" after the
copyright sign. That is true of one competition and of no other, so every
chemistry booklet has been crediting the physics team.

It is `competition.rights_holder` now, and the key is **required** in
`ContextCompetition._schema` rather than optional with a default: a default
would be some other competition's credit, which is the bug that was there.
phys and test say what they always said; chem says "Náboj Chemistry team".

`fks-naboj` is untouched -- its meta matches none of this schema already
(`URL:` rather than `url:`, plus `full:`, `founded:` and a whole `constants:`
table), which is of a piece with nothing building it.

Confirmed on the page: chem/01 prints "©Náboj Chemistry team, 2022." and
phys/29 "©Náboj Physics international team, 2026."

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants